Papers with direct model editing
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing safety mechanisms for Large Language Models (LLMs) are inadequate to protect against jailbreak attacks, resulting in performance degradation on general tasks. |
| Approach: | They propose a method that directly updates a minimal set of relevant parameters to neutralize harmful behaviors while preserving the model’s utility. |
| Outcome: | The proposed model outperforms baseline methods in mitigating jailbreak attacks while preserving the model’s utility. |
Emptying the Ocean with a Spoon: Should We Edit Models? (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study has questioned the use of direct model editing for factual corrections in LLMs. aaron s. de stefano, a sociologist, says that model editing is not a systematic remedy for factuality. |
| Approach: | They argue that direct model editing cannot be trusted as a remedy for LLM disadvantages . authors call for cautious promotion and application of model editing as part of LLM deployment process . |
| Outcome: | The proposed method is not trusted as a remedy for the disadvantages inherent to LLMs, the authors argue . they argue that it opens risks by reinforcing the notion that models can be trusted for factuality . |